Papers by Muhammad Dehan Al Kautsar
IndoSafety: Culturally Grounded Safety for LLMs in Indonesian Languages (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing safety standards are often based on direct translations from English, which overlook key aspects of local communication. |
| Approach: | They propose a high-quality, human-verified safety evaluation dataset tailored for the Indonesian context. |
| Outcome: | The proposed dataset covers formal and colloquial Indonesian, along with three major local languages: Javanese, Sundanese, and Minangkabau. |
What Do Indonesians Really Need from Language Technology? A Nationwide Survey (2025.emnlp-main)
Copied to clipboard
| Challenge: | Despite efforts to develop NLP for Indonesia’s 700+ local languages, progress remains costly due to the need for direct engagement with native speakers. |
| Approach: | They conduct a nationwide survey to assess the actual needs of native Indonesian speakers. |
| Outcome: | The findings indicate that addressing language barriers is the most critical priority . concerns around privacy, bias, and the use of public data highlight the need for greater transparency and clear communication to support broader AI adoption. |
Cultural Benchmarking of LLMs in Standard and Dialectal Arabic Dialogues (2026.acl-long)
Copied to clipboard
Muhammad Dehan Al Kautsar, Saeed Almheiri, Momina Ahsan, Bilal Elbouardi, Younes Samih, Sarfraz Ahmad, Amr Keleg, Omar El Herraoui, Kareem Elzeky, Abed Alhakim Freihat, Mohamed Anwar, Zhuohan Xie, Junhong Liang, Mohammad Rustom Al Nasar, Preslav Nakov, Fajri Koto
| Challenge: | Most benchmarks focus on short text snippets in Modern Standard Arabic (MSA), overlooking cultural nuances that naturally arise in dialogues. |
| Approach: | They propose a culturally grounded conversational dataset covering 13 Arabic-speaking countries, in both Modern Standard Arabic (MSA) and each country’s respective dialect, spanning 12 daily-life topics and 54 fine-grained subtopics. |
| Outcome: | The proposed model performs worse on all three tasks than the MSA benchmark. |